Papers with pretrained sentence encoders
Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations (D19-1)
Copied to clipboard
| Challenge: | Prior work on pretrained sentence embeddings and benchmarks focused on the capabilities of stand-alone sentences. |
| Approach: | They propose a test suite of tasks to evaluate whether sentence representations include broader context information. |
| Outcome: | The proposed training objectives help to encode different aspects of information in document structures. |
Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence Reasoning (2023.eacl-main)
Copied to clipboard
| Challenge: | Recent studies show that textual entailment learning reduces social biases in pretrained sentence encoders. |
| Approach: | They compare pretrained sentence encoders with textual entailment models that learn language logic for downstream language understanding tasks. |
| Outcome: | The proposed models outperform models with lower bias without debiasing processes on stereotype, profession, and emotion bias tests. |
Implicit Discourse Relation Classification: We Need to Talk about Evaluation (2020.acl-main)
Copied to clipboard
| Challenge: | Lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in literature. |
| Approach: | They propose an improved evaluation protocol for implicit relation classification on PDTB 2.0 . they report strong baseline results from pretrained sentence encoders . |
| Outcome: | The proposed evaluation protocol improves the existing framework and provides strong baseline results. |
NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language (P19-1)
Copied to clipboard
| Challenge: | ambiguity in natural language is difficult to interpret due to large linguistic variability. |
| Approach: | They propose to use a Prolog prover to extend neural networks with logic programming to solve multi-hop reasoning tasks over natural language. |
| Outcome: | The proposed model outperforms baseline models on two question answering tasks and is competitive on the MedHop corpus. |